Papers by Anthony Meng Huat Tiong
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches require substantial adaptation of pretrained language models for vision-language reasoning tasks. |
| Approach: | They propose to use natural language and network interpretation as an intermediate representation that glues pretrained models together. |
| Outcome: | The proposed framework outperforms the Flamingo model on VQAv2 and GQA by 8.5%. |